Papers with image and video captioning
A Neural, Interactive-predictive System for Multimodal Sequence to Sequence Tasks (P19-3)
Copied to clipboard
| Challenge: | a neural interactive-predictive system is used to tackle multimodal sequence to sequence tasks . it generates text predictions to different sequence to sequencing tasks, including machine translation, image and video captioning. |
| Approach: | They present a neural interactive-predictive system for tackling multimodal sequence to sequence tasks. |
| Outcome: | The proposed system reduces human effort during the correction process by providing alternative hypotheses. |
MultiCapCLIP: Auto-Encoding Prompts for Zero-Shot Multilingual Visual Captioning (2023.acl-long)
Copied to clipboard
| Challenge: | Existing methods for supervised visual captioning require large scale of images or videos paired with descriptions in a specific language. |
| Approach: | They propose a zero-shot approach that generates captions for different scenarios without labeling . they use concept prompts to retrieve concepts and auto-encode them to learn writing styles . |
| Outcome: | The proposed approach generates captions for different scenarios and languages without labeled vision-caption pairs. |